Back

Clinical Trials

SAGE Publications

Preprints posted in the last 90 days, ranked by how well they match Clinical Trials's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
A Measurement-Based Care Strategy for Buprenorphine-Naloxone Treatment (Bup-MBC): Development of an EHR-Integrated Intervention

Reese, T.; Audet, C.; Ancker, J.; Wright, A.; Marcovitz, D.; Kast, K. A.; Bridges, J.; Tindle, H.; Shah, M.; von Horn, A.; Matheny, M. E.

2026-09-01 addiction medicine 10.64898/2026.08.27.26361539 medRxiv
Top 0.1%
18.6%
Show abstract

Introduction: Risk of recurrent opioid use during buprenorphine-naloxone (bup-nx) treatment is dynamic and remains elevated after initiation, with vulnerability shaped in part by treatment intensity and gaps between visits, yet routine outpatient care relies on episodic encounters and retrospective data. This mismatch can delay recognition of emerging instability and limit timely treatment adjustments. This paper reports the development and specification of an intervention strategy to address this mismatch. Methods: We used a structured, multi-phase design process to specify and configure a measurement-based care (MBC) strategy for bup-nx treatment (Bup-MBC) in outpatient addiction clinics through three phases: (1) a systematic review of patient-reported outcome measures (PROMs) for substance use treatment; (2) a qualitative needs assessment using the Theoretical Domains Framework and COM-B (Capability, Opportunity, Motivation-Behavior) model to identify gaps in risk monitoring, agency, and trust; and (3) iterative co-design with multidisciplinary clinicians to refine workflow fit and trust-preserving use of data. Patients informed item and feedback content during the needs assessment but did not participate in the co-design cycles. Results: Bup-MBC integrates (1) brief between-visit PROMs (e.g., withdrawal, craving, adherence); (2) immediate non-punitive patient feedback; (3) clinician-facing summaries and non-directive prompts in the electronic health record (EHR); and (4) an opt-in between-visit outreach pathway with predefined safety triggers, all configured within existing EHR and patient portal infrastructure. It targets patient and clinician capability to recognize changes in risk, opportunity for action through structured monitoring and visit preparation, and trust and agency through non-punitive communication, without adding substantial burden. The full measure set, severity bands, and question-to-action map are provided as supplementary material. Key trade-offs included prioritizing single-item measures for feasibility, balancing opt-in outreach with safety overrides, and assuming routine clinician use of summaries. Conclusion: This development study specifies an EHR-integrated MBC strategy for outpatient bup-nx treatment. As single-center design work with co-design limited to clinicians and delivery contingent on portal or text-message access, its outputs are hypotheses about mechanism and fit rather than demonstrated effects. Feasibility studies are needed to evaluate uptake, acceptability, workflow fit, and effects on treatment.

2
Comparative Effectiveness of Single vs. Dual WhatsApp Reminders on No-shows: A Target Trial Emulation within the Public Health System of Buenos Aires, Argentina.

Esteban, S.; Quintana, G.; Sanchez, M.; Szmulewicz, A.

2026-08-19 health systems and quality improvement 10.64898/2026.08.17.26360609 medRxiv
Top 0.1%
13.1%
Show abstract

Background: Digital reminders reduce outpatient no-shows, but the optimal timing and frequency of messages remain unclear, particularly in Latin American public health systems. We emulated a target trial to evaluate the comparative effectiveness of four WhatsApp reminder strategies on appointment absenteeism and patient-initiated cancellations. Methods: We analyzed administrative and electronic health-record data from the public health system of the Autonomous City of Buenos Aires, Argentina (June 2023-May 2024). Eligible individuals had scheduled an in-person outpatient appointment in one of 15 prioritized specialties at least 75 hours in advance and had a mobile phone on record. We compared four strategies: (1) dual reminders at ~72 and ~24 hours before the appointment; (2) a single reminder at ~72 hours; (3) a single reminder at ~24 hours; and (4) no reminders. The primary outcome was the proportion of no-shows by the end of follow-up. Secondary outcomes were the cumulative incidence of patient-initiated cancellations overall, within 12 hours of the appointment, and followed by rebooking. We emulated the target trial using a cloning-censoring-weighting approach to estimate per-protocol controlled direct effects, with inverse-probability weights to address time-varying confounding and selection bias. Cumulative incidence of secondary outcomes was estimated using weighted Kaplan-Meier curves. Three pre-specified sensitivity analyses and standardized mean differences assessed robustness and covariate balance. Results: A total of 475,214 first eligible person-appointments were included; baseline no-show risk in the control arm was 34.6%. All three active strategies reduced no-shows compared with no reminders. The single 24-hour reminder produced the largest reduction (Risk Ratio [RR] 0.76, 95% CI 0.72, 0.81; Risk Difference [RD] -8.21 percentage points [pp], 95% CI -9.68, -6.54), followed by the dual-reminder strategy (RR 0.80, 95% CI 0.79,0.81; RD -7.05 pp, 95% CI -7.41, -6.71) and the single 72-hour reminder (RR 0.91, 95% CI 0.84,0.99; RD -3.16 pp, 95% CI -5.69, -0.49). All active strategies increased patient-initiated cancellations relative to control, with the dual-reminder strategy producing the largest increase. Sensitivity analyses preserved the qualitative ranking of strategies across all specifications. Conclusions: In this large target trial emulation, a single just-in-time WhatsApp reminder sent ~24 hours before the appointment was as effective as a dual-reminder schedule in preventing no-shows and superior to a distal 72-hour reminder alone. Adding a second, distal reminder provided no measurable benefit for attendance but substantially increased patient-initiated cancellations, which may be operationally valuable when active slot reallocation is a goal. These findings support timing, rather than frequency, as the primary lever of digital-reminder effectiveness, and favor the deployment of a single proximal reminder as the default strategy in resource-constrained outpatient settings.

3
Automating clinical trial outcome identification and misreporting detection using RegCheck

Cummins, J.; Drysdale, H.; Elson, M.; Hussey, I.; Goldacre, B.; DeVito, N. J.

2026-08-02 epidemiology 10.64898/2026.07.30.26358885 medRxiv
Top 0.1%
9.7%
Show abstract

Objective To evaluate the accuracy and cost of RegCheck, an automated large language model (LLM)-based workflow, for identifying clinical trial outcomes and detecting outcome misreporting by comparing its outputs with manual assessments from the COMPare Trials project. Design Validation study. Setting Sixty-two clinical trials originally assessed in the COMPare Trials project, sampled from five high impact general medical journals. Participants Published clinical trial reports and their corresponding prespecified registrations and/or protocols. Main outcome measures Four prespecified research questions were examined. RQ1 assessed outcome extraction recall relative to COMPare. RQ2 assessed accuracy of outcome classification as primary, secondary, or non-prespecified. RQ3 assessed accuracy of misreporting detection relative to COMPare, with additional manual adjudication of discrepancies between RegCheck and COMPare. RQ4 assessed the average per-paper cost of running the automated workflow. Results Across the validation papers, RegCheck achieved 91.2% outcome extraction recall relative to COMPare, and 83.6% outcome classification accuracy. For detection of outcome misreporting, RegCheck's overall accuracy was 85.6%. However, after resolving discrepancies with the original human judgements (which frequently favoured RegCheck's judgement), revised accuracy for outcome misreporting detection was 94.8%. The mean cost of running the workflow was 5.94 USD per paper. Conclusions RegCheck achieved high overall performance with a rigorous manual benchmark for identifying prespecified and reported outcomes in clinical trials, and detecting outcome misreporting, while operating at very low marginal cost. Adjudication of discrepant judgements suggested that RegCheck frequently identified valid issues not captured in the reference standard. Automated outcome checking may offer a scalable way to support editors, peer reviewers, and authors in detecting outcome switching and improving trial reporting.

4
Insights from a double-blind, randomized, direct-to-participant intervention trial for Long COVID

Vogel, J. M.; Ter Meer, J.; Foster-Bonds, R.; Duff, M. P.; Goosen, A.; Kurakova, A.; Dinh-Luong, E.; Miyasaki, L.; Topol, S.; Sturm, C.; Nowak, C.; Tate, A.; Redd, J.; Shepard, C.; Kheterpal, V.; Steinhubl, S. R.; Topol, E. J.

2026-08-22 infectious diseases 10.64898/2026.08.19.26360832 medRxiv
Top 0.1%
9.7%
Show abstract

Background. Long COVID affects an estimated 400 million people worldwide, and is associated with low quality of life. Nearly all completed Long COVID clinical trials reported no benefit, and most required participants to travel to study sites. This requirement systematically excludes severely affected patients. Because there are numerous candidate therapeutics with established safety profiles and regulatory approvals for other indications, scaled, efficient evaluation of therapeutics is needed. Methods. We designed and are conducting a double-blind, placebo-controlled, phase two trial of tirzepatide for Long COVID fatigue, using an entirely remote infrastructure. Design elements included electronic consent, identity and diagnosis verification through document upload, cold-chain delivery of an injectable study drug through a central pharmacy, shared decision-making for dose titration, repeated at-home capillary blood collection in a biospecimen subcohort, weekly participant touch points through study application, wrist-worn wearable monitoring, and clinical support. The trial is operating under FDA Investigational New Drug authorization. Results. This trial enrolled 1,058 participants in 73 days, at least double the rate of any other Long COVID trial. Mean baseline metrics include mean Fatigue Severity Scale of 59.3 (standard deviation [SD] 4.9), daily step count of 3,611 (SD 2,706, general population reference mean 7,731), EQ-5D-5L of 0.6 (SD 0.2), and FUNCAP27 4.0 (SD 1.0), which was a more severely affected population than other clinical trials that collected comparable data. Study processes are working as designed. Participants use existing advocacy and support channels to gather and communicate. Conclusions. A direct-to-participant, siteless infrastructure can support a double-blind placebo-controlled trial of an injectable drug at scale, accelerate accrual, and reach severely affected participants who are routinely excluded by site-based designs. Modernizing drug distribution and regulatory pathways is needed to realize the full potential of decentralized infrastructure for drug repurposing clinical trials.

5
New tests for trials of very few patients using longitudinal data - a case-study in Autosomal Recessive Cerebellar Ataxias

Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.

2026-09-02 health informatics 10.64898/2026.08.28.26361588 medRxiv
Top 0.1%
6.6%
Show abstract

We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.

6
Rationale and guidance for implementing the continual reassessment method for dose-finding in controlled human infection model studies

Weerasinghe, C.; Osowicki, J.; Simpson, J. A.; Crocker-Buque, T.; McCarthy, J.; Williams, E.; Price, D. J.

2026-07-17 infectious diseases 10.64898/2026.07.16.26358128 medRxiv
Top 0.1%
6.4%
Show abstract

Controlled human infection models (CHIMs) are increasingly used in infectious disease research to study pathogen dynamics and evaluate interventions under controlled conditions. However, these studies are resource-intensive and involve ethical and safety constraints, making efficient study design critical. Dose-finding is a key early component in CHIMs, where the aim is to identify a challenge dose that achieves a target infection probability. Traditional rule-based designs are commonly used but can be inefficient, motivating the use of model-based adaptive approaches such as the Bayesian Continual Reassessment Method (CRM). Although CRM has been extensively studied and widely adopted in Phase I oncology trials for identifying the maximum tolerated dose of therapeutics, its application in CHIM settings remains limited, particularly when the endpoint of interest is infection. This tutorial provides step-by-step guidance for implementing a Bayesian CRM in dose-finding CHIMs, using an oropharyngeal Neisseria gonorrhoeae challenge as a motivating case study. The framework outlines key design components, including dose-grid specification, dose-response model, prior elicitation, Bayesian updating, decision rules, and stopping criteria, with particular emphasis on a clinically interpretable parameterisation. Trial operating characteristics are evaluated through simulation studies under multiple dose-response scenarios and prior-predictive analyses, and compared with a commonly used '3+3' type rule-based design. This work highlights the advantages of Bayesian model-based designs for dose-finding in CHIMs over classic rule-based designs and provides a structured, reproducible framework for implementing CRM, supporting their application in future CHIM studies.

7
Escalate or Switch? Treating the Post-Titration GLP-1 Non-Responder: A Target Trial Emulation With Dose-Equivalence Reclassification

Erly, B.; Raja, S.

2026-07-16 epidemiology 10.64898/2026.07.14.26357491 medRxiv
Top 0.1%
6.2%
Show abstract

Background. When a GLP-1 patient stops responding, should the clinician push the dose or change the drug? Observational answers conflate two distinct sources of confounding. Most early-week "escalations" in real-world data are FDA-mandated titration steps rather than deliberate clinical decisions, and patients who deviate do so for reasons we cannot observe. Semaglutide and tirzepatide are also not equivalent milligram-for-milligram, so naive class-switch comparisons mix mechanism and dose. We resolve both by restricting to post-titration patients and reclassifying treatments under the Whitley 2023 dose-equivalence framework. Methods. From 68,969 telehealth GLP-1 patients we built a post-titration cohort. Each patient's index time is the day they completed at least four weeks at therapeutic dose (Whitley tier 3 or higher: semaglutide 1.0 mg or tirzepatide 5 mg). Confirmed slow response is less than 5% total weight loss at the index, consistent with FDA weight-management drug-development guidance and AACE/ACE criteria. We compared four post-index strategies against continuing the current regimen: within-class dose escalation, equipotent class switch (a Whitley tier change of 1 or fewer), and class switch with potency increase. Direction-specific analyses split switches into semaglutide-to-tirzepatide and tirzepatide-to-semaglutide arms. Outcomes were percent weight loss at 12 and 24 weeks post-index. We estimated effects six ways: propensity-score matching; IPTW with linear and gradient-boosted propensities; the g-formula with linear and gradient-boosted outcome models; and AIPW, the doubly-robust estimator we use as the tiebreaker. Two-layer inverse probability of censoring weighting addressed strategy adherence and outcome ascertainment. We computed E-values, ran a negative-control specification, and stratified by tolerability. Results. The post-titration cohort comprised 24,876 confirmed slow responders. Within-class dose escalation produced a small consistent benefit at 24 weeks: AIPW +0.64 pp (95% CI +0.16 to +1.12), with five non-AIPW estimators ranging +0.47 to +0.76 pp. The continue arm itself lost an additional 8.07 pp over the same window (96% continued to lose), so escalation is a marginal addition to a substantial natural slope, not a rescue. Equipotent class switching from semaglutide to tirzepatide was inconclusive: linear and matching estimators ranged +1.26 to +1.80 pp, but AIPW was -0.33 pp (95% CI -1.27 to +0.60) with only 90 treated patients and limited propensity-score overlap. Class switch with simultaneous potency increase (sema to tirz) gave AIPW +0.65 pp at 12 weeks (95% CI +0.37 to +0.92, n = 80). A negative-control specification yielded ATE -0.13 pp, indicating the pipeline did not generate spurious signal. A held-out-fold prognostic-threshold sensitivity gave a null effect (+0.04 pp), correcting an earlier circular +1.06 pp estimate. Conclusions. Among confirmed post-titration slow responders, within-class dose escalation adds approximately 0.6 percentage points at 24 weeks on top of an 8 percentage point natural slope, consistently across six estimators including doubly-robust inference. This headline effect is small and not robust to modest unmeasured confounding (E-value 1.27) or to MNAR-style outcome attrition (tipping point delta approximately 1.2 pp), so it should be read as hypothesis-generating rather than practice-changing. Class-switching evidence is inconclusive; linear-estimator results suggesting benefit did not survive doubly-robust estimation in small treated samples with limited propensity overlap. The dose-ladder framework, with phase-specific evidence grading, is hypothesis-generating and insufficient on its own to change practice.

8
Mapping the flow of data in clinical trials and assessing its CO2 emissions

Prakasam, H. s.; Mackillop, N.

2026-08-03 health policy 10.64898/2026.07.31.26358778 medRxiv
Top 0.1%
5.5%
Show abstract

Background: Healthcare contributes 5% of global carbon emissions, with academic and industry-sponsored clinical trials forming a meaningful share. Existing trial emission frameworks inadequately capture data storage and analysis emissions, leaving their contribution to overall trial carbon footprint poorly understood. Objective: (i) Map the flow of trial data from the investigator site to the final trial documentation and study report in the trial master files (TMF), and identify emissions hotspots (ii) estimate emissions from the hotspots and assess whether they materially impact overall clinical trial emissions. Design: A top-down assessment of enterprise-level trial data volumes and associated emissions, and trial-level data-related emissions analysis of two representative trials. Results: The data flow mapping highlighted three carbon emission hotspots: (a) data analysis in a statistical environment (e.g., entimICE) (b) Trial Master File (TMF) storage, and (c) short-term and long-term data storage by external clinical research organisations (CROs). Within entimICE, the total volume of trial data stored for analysis across all active and recently completed trials at AZ was 100-125TB, stored across four servers in Sweden, generating 80 to 100 tonnes CO2eq annually. TMF storage emissions fell below measurable thresholds. CROs stored substantial data volumes, but per-trial emissions were likely insignificant due to economies of scale in large data centres. Conclusion: To our knowledge, this is the first study to comprehensively assess the carbon footprint of industry-sponsored clinical trial data storage and analysis. The study highlighted the complex network of nodes and junctions involved in managing trial data. The results suggest that carbon emissions from trial data storage and processing are currently a small proportion and unlikely to materially impact overall trial-related emissions. Future research should confirm these results with clinical trials that employ emerging data-intensive computational operations, such as the integration of artificial intelligence (AI).

9
Clinical outcomes of HIV treatment clients enrolled in six-month dispensing after less than 6 months on treatment in Zambia: A target trial emulation

KACHINGWE, E.; Fox, M. P.; Ntjikelane, V.; Mokhele, I.; Shumba, K.; Rosen, S.; Kamanga, A.; Haimbe, P.; Sivile, S.; Huber, A. N.

2026-08-02 hiv aids 10.64898/2026.07.30.26358576 medRxiv
Top 0.1%
5.4%
Show abstract

Background: Six-month multi-month dispensing (6MMD) of antiretroviral therapy (ART) reduces clinic visit frequency and is associated with improved retention in care. During the COVID-19 pandemic, Zambia offered 6MMD to clients 3 months after ART initiation, rather than 6 months standard requirement. We estimated the effect of early (3<6 months on ART) versus standard (6-12 months) 6MMD enrolment on the rate of treatment interruption. Methods: We emulated a target trial using routinely collected electronic medical records from 12 public health facilities in Zambia. Eligible clients were 15 years and above, initiated ART 01/20-08/22, were WHO stage 1 or 2 at ART initiation, and had more than 21 months of potential follow-up. Treatment interruption was defined as missing a scheduled clinic or pharmacy visit by more than 28 days. We applied a clone-censor-weight approach to reduce immortal time bias. Clones were censored when observed dispensing deviated from their assigned strategy. Inverse probability of censoring weights (IPCW) accounted for informative censoring, while inverse probability of treatment weights (IPTW) balanced measured baseline confounders between strategies. We used weighted pooled logistic regression of person-month data to estimate the odds of treatment interruption between early and standard 6MMD enrollers, including follow-up months to model the monthly baseline risk. Results: A total of 6,142 ART clients met the inclusion criteria. 741 (12.1%) were early 6MMD enrollers, 1,590 (26.1%) standard 6MMD enrollers, and 3,811 (62.0%) eligible clients who never enrolled in 6MMD. During follow-up, 268 treatment interruptions occurred. In the primary analysis, early 6MMD was associated with lower odds of treatment interruption than standard 6MMD OR 0.701 (95% CI 0.51-0.97). The predicted cumulative probability of treatment interruption at 18 months was 6.5% under the early 6MMD strategy and 9.1% under the standard strategy (risk difference: -2.6 percentage points). Conclusions: Enrolment in 6MMD at 3-6 months after ART initiation was associated with lower odds of treatment interruption than standard enrolment at 6-12 months, with a predicted absolute risk difference of -2.6 percentage points at 18 months. We found no evidence that earlier access to 6MMD increases the risk of treatment interruption.

10
Forecasting U.S. 12-Month-Ending Overdose Death Counts: Multi-Model National and Regional Projections

Alhassan, F.; Karami, H.; Bohler, R.; Fung, I. C.-H.; Mamelund, S.-E.; Lee, S.; Peterson, E.; Chowell, G.

2026-08-05 addiction medicine 10.64898/2026.08.03.26359533 medRxiv
Top 0.1%
4.4%
Show abstract

Aims: To assess whether recent declines in U.S. rolling 12-month-ending drug overdose death counts are projected to continue, compare the retrospective performance of short-term forecasting models, and estimate national and regional 12-month changes. Design: Comparative time-series forecasting study with a retrospective March 2025-February 2026 forecast evaluation and subsequent 12-month-ahead projections through February 2027. Setting: United States and four U.S. Census regions: Northeast, Midwest, South, and West. Cases: Aggregate drug overdose deaths reported in the National Center for Health Statistics Vital Statistics Rapid Release system (VSRR) and identified using ICD-10 underlying cause-of-death codes X40-X44, X60-X64, X85, and Y10-Y14. Measurements: The primary outcome was the monthly series of rolling 12-month-ending overdose deaths. Candidate models included ARIMA, generalized additive models, Prophet, and AICc-ranked n-sub-epidemic models. Models were calibrated using January 2020-February 2025 data and evaluated against March 2025-February 2026 observations using mean absolute error, mean squared error, empirical 95% prediction interval coverage, and weighted interval score (WIS). Individual models were ranked by retrospective WIS, and normalized inverse-WIS weights were used to construct ensembles from the top-ranked models. Final forecasts were generated for March 2026-February 2027 after recalibrating models using January 2020-February 2026 data. Results: Retrospective WIS performance differed geographically: GAM performed best nationally and in the Midwest and West, the leading n-sub-epidemic model in the Northeast, and ARIMA in the South. Median forecasts from all individual models and ensembles projected declines from February 2026 to February 2027 nationally and in each region, although prediction intervals varied substantially. Individual national median projections ranged from declines of 12.8% to 25.8%, while weighted-ensemble median projections indicated declines of 20.5% to 21.7%. Projected weighted-ensemble declines were larger in the Northeast (25.4%-30.0%) and Midwest (24.3%-26.3%) than in the South (16.6%-19.5%) and West (18.0%-21.8%). Summary: Ensemble forecasts of the reported provisional VSRR series were consistent with continued declines in rolling annual overdose death counts nationally and across U.S. Census regions. The projected magnitude of decline differed by region and model specification. Because the forecasts used provisional rolling 12-month-ending counts, they should be interpreted as surveillance projections rather than exact monthly mortality predictions.

11
A PRISMA-Aligned Agentic Framework for Medical Systematic Reviews and Evidence Synthesis

Huang, H.; Zheng, Q.; Qiu, P.; Zhao, W.; Zhang, Y.; Xie, W.; Wang, Y.; Zhang, X.; Wu, C.

2026-08-02 health informatics 10.64898/2026.07.30.26359375 medRxiv
Top 0.1%
4.3%
Show abstract

Medical systematic reviews are central to evidence-based medicine, but they remain slow, labor-intensive, and difficult to maintain under the full Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) workflow. Recent LLM-based deep research agents offer a promising route to addressing this challenge, yet reliable deployment in medical systematic reviews remains limited by insufficient clinical domain knowledge and inconsistent adherence to evidence-based methodological standards across the full workflow. We address these gaps with MedSR-Copilot, a PRISMA-aligned multi-agent copilot that decomposes review automation into literature retrieval, coarse-to-fine screening, data extraction, Risk-of-Bias assessment, and evidence synthesis, while preserving structured intermediate artifacts throughout the workflow. We further introduce MedSR-Bench, an end-to-end benchmark for evaluating systems beyond isolated subtasks, from review input to final evidence-synthesis conclusions. MedSR-Copilot completes medical systematic reviews end-to-end under the full PRISMA workflow, achieving 63.6% human-aligned conclusions, 18.3 percentage points above the best baseline among strong general-purpose LLMs and prior automated review systems. In a human-AI collaboration study involving 23 analysis groups across four systematic review topics, MedSR-Copilot, used as a copilot, reduces end-to-end review time by 64.9% and improves final conclusion accuracy by 27.4 percentage points compared with routine-practice workflows. Together, these results demonstrate the reliability and efficiency of MedSR-Copilot as a medical research copilot and suggest a practical path toward trustworthy review automation.

12
Autonomous generation of decision-grade clinical evidence

Yang, S.; Wu, J.; Xie, H.; Xin, Z.; Wang, W.

2026-07-01 health informatics 10.64898/2026.06.26.26356653 medRxiv
Top 0.1%
4.2%
Show abstract

Medical practice is bottlenecked by the slow production of high-quality clinical evidence. Despite progress in automating selected stages, autonomous conduct of the entire research life cycle remains beyond reach. Here we present OpenEBM, the first autonomous system to generate decision-grade clinical evidence by conducting evidence-synthesis research end to end. To enable and evaluate this, we develop OpenEBM-Corpus, a foundation resource of expert-annotated research trajectories that enables training of a specialist model, and OpenEBM-Bench, a multidisciplinary benchmark that evaluates the entire research life cycle. Our compact specialist model generates valid clinical evidence in 90.7% of end-to-end evaluations and matches expert performance across the research trajectory, whereas GPT-5 falls to 3.8% as failures propagate through dependent stages. In blinded evaluations across clinical domains, independent evaluators prefer OpenEBM at multiple stages and cannot distinguish its reasoning traces from expert-conducted work above chance. Applied to a question left unresolved by current guidelines, OpenEBM produces de novo evidence addressing the efficacy and safety of neoadjuvant chemotherapy for locally advanced rectal cancer. OpenEBM brings within reach the founding aspiration of evidence-based medicine and establishes a paradigm for scalable evidence generation.

13
Towards understanding the disease landscape of clinical trials in Germany: Ontology and embedding-based pipelines versus Large Language Models for ICD-10 Harmonization

Ndabashinze, R.; Franzen, D.; Kozuch, E.; Aagerup, J.; Fink, A.; Yerunkar, S. S.; Hunter, K.; Mayo-Wilson, E.; Ying, X.; Kilicoglu, H.; Schorr, S. G.; Seidler, A. L.

2026-08-06 health informatics 10.64898/2026.08.04.26359616 medRxiv
Top 0.1%
3.9%
Show abstract

Background Clinical trials conducted in Germany are registered across multiple registries, including the German Clinical Trials Register (DRKS), ClinicalTrials.gov, the EU Clinical Trials Register (EUCTR), and, since 2023, the Clinical Trials Information System (CTIS). These registries record health conditions using different classification systems and terminologies, including ICD-10-GM, MeSH, MedDRA, and free text, making cross-registry analyses difficult. We developed and evaluated a pipeline for harmonizing trial condition descriptions to WHO ICD-10 and compared its performance with that of a large language model (LLM) and to health conditions coded by humans. Methods We developed a four-stage, registry-aware mapping pipeline consisting of: (i) condition mention extraction and normalization; (ii) classification of ICD-mappable versus non-mappable mentions; (iii) ontology-based candidate generation using UMLS links between MeSH, MedDRA, ICD-10-GM, and WHO ICD-10; and (iv) SapBERT-based semantic retrieval with hybrid confidence scoring. A second variant additionally applied cross-encoder reranking of the top candidate codes. A stratified sample of 500 condition mentions was manually coded to create an expert reference standard. GPT-4o was evaluated in parallel using the same structured decision framework as the human reviewers. Performance was assessed using accuracy, precision, F1 score, and Cohen's {kappa} at the three-character, block, and chapter levels of ICD-10. Results The pipeline was applied to 23,061 clinical trials and identified 39,512 ICD-mappable condition mentions, of which 72.4% received a high-confidence assignment. Against 390 expert-coded mentions, the baseline pipeline achieved 49.0% accuracy at the three-character ICD-10 level ({kappa} = 0.487), increasing to 58.7% at the chapter level ({kappa} = 0.561). The cross-encoder method produced small but consistent improvements across all evaluation levels. Candidate-recall analysis showed that the correct code was present in the retrieved candidate set in only 73.7% of cases. The LLM substantially outperformed both pipeline variants, achieving 96.7% accuracy and near-perfect agreement with expert coding ({kappa} = 0.966) at the three-character level. The LLM also assigned clinically plausible codes to 82.4% of rejected mentions, 62.8% of Tier-3 exclusions, and 92.3% of review-band mentions. Conclusion Automated harmonization of clinical trial condition data across heterogeneous registries is feasible and supports the use of a common ICD-10 framework for cross-registry analyses. The LLMs achieved high agreement with expert coding, and performed better than the deterministic ontology and embedding pipeline, which achieved moderate agreement. These findings indicate that LLMs can support analyses of the distribution of health conditions investigated in clinical trials in Germany.They are a promising tool for classification of other non-standardised trial characteristics in registries. Keywords: Clinical trial registries; ICD-10; disease harmonization; UMLS; entity linking; SapBERT; large language models; clinical research; natural language processing.

14
Balancing Relapse Risk and Agency in Buprenorphine-naloxone Treatment: A Qualitative Needs Assessment to Inform Patient-Centered Care

Reese, T.; Shah, M. V.; Wright, A.; Matheny, M. E.; Marcovitz, D. E.; Kast, K. A.; Bridges, J.; Tindle, H.; von Horn, A.; Audet, C.

2026-08-23 addiction medicine 10.64898/2026.08.21.26360804 medRxiv
Top 0.1%
3.7%
Show abstract

Objectives Outpatient buprenorphine-naltrexone (bup-nx) treatment reduces overdose risk, yet many patients still return to use or disengage from treatment. We sought to understand how patients and prescribers experience and manage relapse risk, monitoring, and treatment agency in routine bup-nx treatment to identify gaps in current practice. Methods We conducted a qualitative needs assessment using semi structured, critical incident interviews with patients receiving outpatient bup-nx and prescribers who manage bup-nx treatment. Interviews examined situations involving relapse risk and empowerment in treatment decisions. We structured data collection and analysis using the Theoretical Domains Framework and COM B model to characterize determinants. Transcripts were coded deductively and inductively until code level saturation was reached. Results Participants (9 patients, 8 prescribers) described nine treatment needs mapped to the Capability, Opportunity, and Motivation components of the COM B model. These themes highlighted how patient agency in bup-nx treatment was constrained by physiologic and emotional states, with withdrawal, craving, pain, and distress often overriding longer term goals. Relapse vulnerability was experienced as dynamic and intensifying between visits, while clinical detection remained anchored to visit bound assessments, urine drug testing, refill patterns, and crisis driven contact, creating blind spots. Structural friction (pharmacy rules, insurance disruptions, transportation and housing instability), stigma from family and recovery communities, and motivational processes tied to fluctuating readiness and trust in monitoring further shaped engagement, disclosure, and dosing decisions; the same monitoring tools could either support honest disclosure or provoke concealment when perceived as punitive. Conclusions Relapse risk and agency in bup-nx treatment are negotiated as dynamic processes within structurally constrained and trust sensitive systems. Addressing the identified capability, opportunity, and motivation gaps will require patient centered, trust preserving approaches to monitoring and shared decision making.

15
Assessment of Zero-Shot Large Language Model (LLM) Assisted Clinical Trial Matching Processes: A Metastatic Cancer Use Case

Weng, Y.; Yalamaddi, H.; Fu, D.; Mishra, A.; Bunning, B. J.; Martin, A. B.; Hope, J.; Charu, V.; Kurian, A.; Desai, M.

2026-07-10 oncology 10.64898/2026.07.06.26354647 medRxiv
Top 0.1%
3.6%
Show abstract

Introduction: For oncology patients with limited treatment options, clinical trials may be a critical lifesaving pathway. Identifying relevant trials, however, is a time-consuming and difficult task. Several patient-trial matching processes incorporating large language models (LLMs) have been proposed to alleviate the burden on patients and oncologists. We aim to explore the benefits and practical challenges of zero-shot LLM-assisted trial matching processes by analyzing the results for a single pancreatic cancer patient. Materials and Methods: The results of a simple zero-shot LLM-assisted clinical trial matching process for our patient were compared to those of a "human benchmark," which was developed manually by two of the authors interfacing directly with ClinicalTrials.gov. Performance metrics -- sensitivity, specificity, precision, and accuracy -- were calculated. In addition, a qualitative content analysis (QCA) of LLM reasoning text was done to identify patterns in "errors," which we define as a human-LLM discrepancy in final patient eligibility. Implications and severity of errors are discussed. Results: The zero-shot LLM-assisted process returned potential trials with a sensitivity, specificity, and precision of 81.1%, 89.3%, and 86.5% respectively compared to the human benchmark. Qualitative error analyses revealed that about 73% of errors could potentially be alleviated with improved prompting and information access. Overall performance seemed comparable to that of human reviewers. Conclusion: The results from this preliminary real-world case study provide additional evidence to the literature in support of the integration of LLMs in clinical trial matching to provide benefit to patients with metastatic cancer with limited options.

16
Simulation-Guided Selection of a Bayesian Adaptive Phase II Design for a Nine-Arm Cilostazol-Albumin Trial in Aneurysmal Subarachnoid Hemorrhage

Qureshi, A. I.; Raza, H.; Alam, N.; Beall, J.; Gajewski, B. J.; Martin, R. L.; Suarez, J. I.

2026-06-22 neurology 10.64898/2026.06.18.26356019 medRxiv
Top 0.1%
3.5%
Show abstract

Background: The Cilostazol Albumin Treatment in Subarachnoid Hemorrhage (CATS) trial evaluates eight active cilostazol-human albumin regimens plus control in patients with aneurysmal subarachnoid hemorrhage. We summarized the rationale for the primary statistical design, compared alternative Phase II methodologies, and evaluated reduced-arm sensitivity scenarios. Methods: The binary primary endpoint is Common Data Elements-defined delayed cerebral ischemia within 14 days after randomization. The selected design is Bayesian adaptive, with a burn-in phase, response-adaptive randomization among active arms while maintaining fixed control allocation, four interim analyses, early stopping for expected success or futility, and a two-dimensional normal dynamic linear model. Primary operating characteristics were obtained from 1,000 virtual trials per scenario using Fixed and Adaptive Clinical Trial Simulator version 7.0.0. Exploratory simulations evaluated six-, four-, and two-active-arm configurations and simplified alternative designs. Results: Compared with fixed equal allocation, the Bayesian adaptive design preserved an approximately 10% false-success probability under the global null while improving probability of success and efficiency in clinically relevant scenarios. Under the Realistic scenario, probability of success increased from 0.61 to 0.86, expected sample size decreased from 400 to 308, and expected duration decreased from 235 to 187 weeks. Under common thresholds, null probability of success was 0.098 for the full anchor and 0.073 for Reduced-6; Reduced-6 probabilities of success were 0.774 and 0.765 in the Realistic and Realistic2 scenarios. However, Reduced-6 omitted two monotherapy anchors and was less robust in Backwards2. In the comparator simulation, the selected design had probability of success of 0.858 and expected sample size of 308.3 under the Realistic scenario, compared with 0.624 to 0.845 and approximately 352 to 400 for simplified comparators. Conclusions: For identifying the most promising cilostazol-human albumin regimen for Phase III rather than confirming efficacy, the Bayesian response-adaptive design with two-dimensional normal dynamic linear model borrowing is more efficient and better aligned than simplified comparators. The full nine-arm design remains preferable because it preserves the complete therapeutic discovery space and is more robust to misspecified or non-smooth response surfaces.

17
From Screening to Sustained Recovery: A Multidomain Systematic Review and Evidence Map of Adolescent Substance-Use Rehabilitation with Nested Meta-analysis of Youth Opioid Treatment

Mittal, P.; Srivastava, A.; Singh, P. P.; Chauhan, J.

2026-07-13 addiction medicine 10.64898/2026.07.11.26357831 medRxiv
Top 0.1%
3.4%
Show abstract

Background: Adolescent substance-use rehabilitation is a care-continuum problem spanning detection, engagement, active treatment, relapse prevention, aftercare, family support, and equity-oriented implementation. Existing reviews are often modality-specific and do not show how evidence aligns with substances, populations, outcomes, stages of care, or policy needs. Objectives: To map and synthesise the 2015-2025 adolescent and transitional-age youth SUD rehabilitation literature across intervention domains, stages, substances, outcomes, equity/disadvantage, geography, and economics, and to perform meta-analysis only where pooling was clinically defensible. Methods: PubMed, Scopus, and Web of Science records were harmonised to 2015-2025 and deduplicated. Two reviewer roles applied a predefined charting codebook for substance focus, technique family, rehabilitation stage, equity/disadvantage flags, outcome family, and study-design signal. Evidence was synthesised across AI/digital, psychiatric/psychotherapeutic, pharmacological, family/social, behavioural, residential/continuing-care, school/community, harm-reduction, and policy domains. Random-effects meta-analysis was restricted to comparative youth OUD medication-supported trials with extractable binary outcomes. Results: The search identified 1,676 records; 554 duplicates were removed, leaving 1,122 unique records. Metadata screening retained 579 records for evidence-map charting: 112 high-confidence records and 467 conservative metadata-supported records requiring full-text verification before final selective-journal submission. The charted evidence was concentrated in active treatment (n=433) and relapse prevention (n=114); aftercare/follow-up was weak (n=8). Intervention-family signals were led by pharmacological/MOUD (n=72), psychotherapy/psychiatric care (n=65), school/community/brief interventions (n=46), residential/continuing care (n=41), family/social therapy (n=30), AI/digital/telehealth (n=25), harm-reduction/policy (n=24), and CM (n=22). The primary youth OUD retention/completion meta-analysis favoured medication-supported treatment (OR 7.67, 95% CI 3.98-14.78; I^2=0%; k=2; n=188). An exploratory favourable-outcome analysis produced a similar estimate (OR 7.94, 95% CI 4.24-14.89; I^2=0%; k=3; n=229). Conclusions: The strongest pooled quantitative claim supports medication-supported treatment for youth OUD. For non-opioid substances, digital care, family therapy, CM, residential care, aftercare, and equity-oriented implementation, the literature is clinically important but not yet consistently synthesis-ready. Future trials should evaluate complete care pathways, adopt core outcomes, report age-banded and equity subgroup effects, and include economic and implementation endpoints.

18
Bayesian Borrowing of External Information in Clinical Trials: A Comparison of MAP, RMAP, and SAM Priors

Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.

2026-08-31 pharmacology and therapeutics 10.64898/2026.08.26.26360843 medRxiv
Top 0.1%
3.3%
Show abstract

Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.

19
The quality-failure paradox in extracellular vesicle therapeutics: quantitative benchmarking of 152 clinical trials against CAR-T cell therapies.

Glebov, K.; Zarovni, N.

2026-07-22 health economics 10.64898/2026.07.20.26358510 medRxiv
Top 0.1%
3.3%
Show abstract

The stall in clinical translation of extracellular vesicle (EV) therapeutics requires addressing critical and hidden causes of attrition and delays. Despite more than 150 registered clinical trials and fifteen years of clinical investigation, small EV (sEV) therapeutics have yielded zero FDA approvals, a translational deficit that has received remarkably little quantitative scrutiny. We systematically evaluated 783 EV-related clinical trials (152 of which therapeutic) registered through December 2025, benchmarking the sEV pipeline against 1,131 CAR-T cell therapy trials (ex-China) that share the "process is the product" manufacturing constraint, yet have delivered six FDA-approved products. A quality-failure paradox emerges, in which industry-sponsored sEV trials exhibit 39% higher composite methodological quality (quality index 2.49 vs. 1.79) yet fail at approximately five-fold the rate of academic programmes (28.6% vs. 5.1%; OR=7.40, 95% CI 2.5-22.3, p=0.0004). Cross-modality replication in CAR-T trials confirms the direction of this association (OR=2.51, p<0.001; Cochran-Mantel-Haenszel pooled OR=2.76, p<0.001). Registry abandonment, affecting 24% of academic sEV trials, constitutes a hidden failure mode that, when reclassified, dissolves the apparent paradox. Premature clinical entry of incompletely defined products, rather than insufficient methodological rigour, represents the central constraint on sEV translational progress. The opacity of research data, inadequate clinical evaluation, and failure to report negative (null) findings in clinical trials create a hidden crisis that drains hundreds of millions invested by public and private sector-eroding the return on investment-and traps progress in cycles of duplicated effort, wasteful resource allocation, and missed learning opportunities.

20
SymPerturb converts symptom-network structure into testable intervention priorities

Zhu, Z.; Yu, J.; Hu, T.; Yang, Z.; Wang, J.

2026-07-30 epidemiology 10.64898/2026.07.27.26359002 medRxiv
Top 0.1%
3.2%
Show abstract

Symptom networks encode conditional dependence but do not by themselves identify causal or clinically actionable intervention targets. We introduce SymPerturb, a virtual-perturbation framework that distinguishes four primitive perturbation operators - virtual knockout, virtual knockdown, edge-level communication blocking and node-centred communication blocking - from three analytic procedures - virtual dosage perturbation, combination perturbation and sequence optimisation. The reference Gaussian implementation is embedded in a general location-scale map with symptom-specific target anchors, making explicit that zero anchoring and linked mean-variance attenuation are modelling choices. Seven utility outcomes quantify downstream efficacy, dose efficiency, breadth, cross-module reach, communication blocking, combination value and responsiveness; robustness is reported separately as an uncertainty diagnostic. Their direction-aligned, within-candidate-set weighted mean defines the virtual perturbation priority score (VPPS), which is a relative ranking rather than a transportable clinical utility score. In a known 22-node, four-module generating network, analytical efficacy agreed with 100,000-draw Monte Carlo estimates within 0.0024 standard deviations. The reported finite-sample VPPS results were generated with the original eight-component exploratory score and therefore require regeneration under the revised seven-utility-dimension definition. These simulations provide internal computational verification under model compatibility, not causal or external validation. SymPerturb is intended to generate auditable target hypotheses for longitudinal and experimental testing.